Checkpointing Strategies for Scheduling Computational Workflows
نویسندگان
چکیده
منابع مشابه
Checkpointing Strategies for Scheduling Computational Workflows
We study the scheduling of computational workflows on compute resources that experience exponentially distributed failures. When a failure occurs, rollback and recovery is used to resume the execution from the last checkpointed state. The scheduling problem is to minimize the expected execution time by deciding in which order to execute the tasks in the workflow and deciding for each task wheth...
متن کاملCheckpointing Strategies for Scheduling Computational Workflows Guillaume
We study the scheduling of computational workflows on compute resources that experience exponentially distributed failures. When a failure occurs, rollback and recovery is used to resume the execution from the last checkpointed state. The scheduling problem is to minimize the expected execution time by deciding in which order to execute the tasks in the workflow and deciding for each task wheth...
متن کاملCheckpointing Based Fault Tolerant Job Scheduling System for Computational Grid
A computational grid environment, due to its heterogeneous, autonomous and dynamic nature is prone to different kinds of faults which may lead to delay in completion of job or even execution of job from starting point. Checkpointing mechanism plays a vital role for making grid more reliable, cost effective and efficient. In this paper, we have proposed schemes based on system checkpointing and ...
متن کاملDynamic computational workflows: Discovery, optimisation and scheduling
The Grid computing community is converging on a service-oriented architecture in which applications are composed from geographically-distributed, interacting web services, and are expressed in a workflow description language, usually based on XML. Such workflows are viewed as offering a useful representation of service-based applications or applications composed of standalone components that ar...
متن کاملCombining Checkpointing and Replication for Reliable Execution of Linear Workflows
This report combines checkpointing and replication for the reliable execution of linear workflows. While both methods have been studied separately, their combination has not yet been investigated despite its promising potential to minimize the execution time of linear workflows in failure-prone environments. The combination raises new problems: for each task, we have to decide whether to checkp...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
ژورنال
عنوان ژورنال: International Journal of Networking and Computing
سال: 2016
ISSN: 2185-2839,2185-2847
DOI: 10.15803/ijnc.6.1_2